Papers with cross-attention mechanism

12 papers
Dual-Encoder Transformers with Cross-modal Alignment for Multimodal Aspect-based Sentiment Analysis (2022.aacl-main)

Copied to clipboard

Challenge: Multimodal aspect-based sentiment analysis (MABSA) aims to extract aspect terms from text and image pairs, and then analyze their corresponding sentiment.
Approach: They propose a dual-encoder transformer with cross-modal alignment to extract aspect terms from text and image pairs and then analyze their corresponding sentiments.
Outcome: The proposed approach outperforms existing methods on two benchmarks.
TrendPulse: A Simple yet Efficient Framework for Capturing Viral E-Commerce Spikes via LLM-Driven Contextualization (2026.acl-industry)

Copied to clipboard

Challenge: Modern e-commerce platforms mostly depend on reactive discovery, where products surface only after users search for them.
Approach: They propose a framework that identifies regional search momentum and leverages Large Language Model to transform spikes into semantic trends.
Outcome: The proposed framework shows consistent improvements across multiple business metrics and overall user experience.
M3: A Multi-View Fusion and Multi-Decoding Network for Multi-Document Reading Comprehension (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods for multi-document reading comprehension cannot make full of the advantages of both approaches.
Approach: They propose a multi-view fusion and multi-decoding method that integrates multiple documents for answering questions.
Outcome: The proposed method improves on two mainstream multi-document reading comprehension datasets.
Context-Aware Interaction Network for Question Matching (2021.emnlp-main)

Copied to clipboard

Challenge: Existing models focus on word-level local matching and neglect the importance of contextual information.
Approach: They propose a context-aware interaction network to properly align two sequences and infer their semantic relationship by using gate fusion layers.
Outcome: The proposed model can accurately align two sequences and infer their semantic relationship on two question matching datasets.
Dior-CVAE: Pre-trained Language Models and Diffusion Priors for Variational Dialog Generation (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing variational dialog models have pre-trained, restricting diversity of responses . a diffusion model increases complexity of prior distribution and its compatibility with PLMs .
Approach: They propose a hierarchical conditional variational autoencoder with diffusion priors to address these challenges.
Outcome: The proposed method generates more diverse responses without dialog pre-training.
EtriCA: Event-Triggered Context-Aware Story Generation Augmented by Cross Attention (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for story generation still suffer from problems of relevance and coherence.
Approach: They propose a novel neural generation model which maps contextual and event features to event sequences with a cross-attention mechanism and exploits logical relatedness between events.
Outcome: The proposed model outperforms state-of-the-art models on automatic and human evaluations and shows that it can leverage contextual and event features.
Large Sequence Representation Learning via Multi-Stage Latent Transformers (2022.coling-1)

Copied to clipboard

Challenge: a novel algorithm for named-entity recognition (NER) uses language and spatial features to predict entity tags for structured text . a dataset of 11,926 images depicting food product labels is used to perform NER tasks .
Approach: They propose a multi-stage transformer architecture for named-entity recognition . they propose RADAR, an LSTM classifier operating at character level, to refine NER predictions .
Outcome: The proposed method outperforms two competing models on a food label dataset.
Improving Retrieval Augmented Open-Domain Question-Answering with Vectorized Contexts (2024.findings-acl)

Copied to clipboard

Challenge: Retrieval Augmented Generation can be used to process long contexts in Open-Domain Question-Answering tasks.
Approach: They propose a method to cover longer contexts in Open-Domain Question-Answering tasks by using a small encoder language model and cross-attention with origin inputs.
Outcome: The proposed method can cover longer contexts while keeping the computing requirements close to the baseline.
Graph-to-Text Generation with Dynamic Structure Pruning (2022.coling-1)

Copied to clipboard

Challenge: Recent studies show that explicitly modeling the input graph structure can significantly improve the performance.
Approach: They propose a structure-aware cross-attention mechanism to re-encode the graph representation conditioning on the newly generated context at each decoding step.
Outcome: The proposed model improves performance on two graph-to-text datasets with only minor increase on computational cost.
Query-based Cross-Modal Projector Bolstering Mamba Multimodal LLM (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing models that capture multiple modalities with a single input length are unable to handle this computational burden.
Approach: They propose a query-based cross-modal projector that compresses visual tokens based on input through the cross-attention mechanism.
Outcome: The proposed projector reduces the need for manually designing the 2D scan order of original image features when converting them into an input sequence for Mamba LLMs.
Fine-tuning LLMs with Cross-Attention-based Weight Decay for Bias Mitigation (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) excel in natural language processing tasks but often propagate societal biases from their training data, leading to discriminatory outputs.
Approach: They propose a method that modifies the LLM architecture to mitigate bias by adjusting the attention weights of sensitive tokens.
Outcome: The proposed method can handle multiple sensitive attributes and does not require full knowledge of sensitive tokens presented in the dataset.
NLoPT: N-gram Enhanced Low-Rank Task Adaptive Pre-training for Efficient Language Model Adaption (2024.lrec-main)

Copied to clipboard

Challenge: Pre-trained Language Models (PLMs) have superior performance on downstream tasks . however, conventional TAPT adjusts all parameters of the PLMs, which distorts the learned generic knowledge embedded in the original PLM's weights.
Approach: They propose a two-step n-gram enhanced low-rank task adaptive pre-training method to customize a PLM to the downstream task.
Outcome: The proposed method improves performance on six datasets from four domains.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations